Data Engineer

Data Engineering United States


Description

Key responsibilities 
• Build ingestion into the bronze layer for assigned sources: gateway and observability logs, productivity 
tool admin APIs, AI-enabled SaaS usage, hyperscaler billing exports and reference data. Land raw and 
untransformed, on a scheduled refresh, replayable if the downstream design changes. 
• Work to the shared bronze landing contract so each tool is ingested once and serves both this program 
and the parallel productivity initiative, rather than being integrated twice. 
• Build the silver layer: typed, deduplicated and conformed to the canonical dimensions, refreshed 
independently of any downstream publication schedule. 
• Build gold marts carrying attribution method, attribution level, cost basis and provisional status alongside 
cost and usage. 
• Implement the attribution and allocation logic designed by the analysts, including precedence resolution 
and ratio-based splitting of shared endpoint cost. 
• Work within Unity Catalog governance — shared bronze and silver, separate gold marts with a recorded 
owner per dataset — including permissions, lineage and cataloging. 
• Implement data quality rules and monitoring: completeness, freshness and tag-coverage checks with 
alerting, so pipeline problems surface before they reach a divisional invoice. 
• Manage the volume impact of enabling caller-identity data in the cost and usage report, which multiplies 
row counts by the number of calling identities per model. 
• Work to the per-source cadence — daily where controls and anomaly detection depend on it, monthly 
where they do not — within the team's existing CI/CD and promotion practices. 
Essential skills and experience 
• Advanced Databricks engineering: Delta Lake, medallion architecture, Databricks Workflows, Auto 
Loader and incremental ingestion patterns. 
• Unity Catalog to a governance standard — catalogs, schemas, permissions, lineage — not merely as a 
place tables happen to live. 
• Strong Python and PySpark, and strong SQL. Notebook-based development. 
• Ingestion from REST APIs including pagination, throttling, incremental watermarks and credential 
handling, plus cloud object storage across AWS, Azure and GCP. 
• Performance and cost optimization of Spark workloads: partitioning, clustering, file sizing and cluster 
configuration. 
Tokenomics Program - Contract Role Descriptions  |  Page 7 
• CI/CD for Databricks — asset bundles or equivalent — and Git-based development workflow. 
• Able to work to an existing catalog structure and coding standard rather than introducing a parallel 
approach.